Additive Logistic Regression : a Statistical View ofBoostingJerome Friedman
نویسندگان
چکیده
Boosting (Freund & Schapire 1996, Schapire & Singer 1998) is one of the most important recent developments in classiication methodology. The performance of many classiication algorithms can often be dramatically improved by sequentially applying them to reweighted versions of the input data, and taking a weighted majority vote of the sequence of classiiers thereby produced. We show that this seemingly mysterious phenomenon can be understood in terms of well known statistical principles, namely additive modeling and maximum likelihood. For the two-class problem, boosting can be viewed as an approximation to additive modeling on the logistic scale using maximum Bernoulli likelihood as a criterion. We develop more direct approximations and show that they exhibit nearly identical results to boosting. Direct multi-class generalizations based on multinomial likelihood are derived that exhibit performance comparable to other recently proposed multi-class generalizations of boosting in most situations, and far superior in some. We suggest a minor modiication to boosting that can reduce computation, often by factors of 10 to 50. Finally, we apply these insights to produce an alternative formulation of boosting decision trees. This approach, based on best-rst truncated tree induction , often leads to better performance, and can provide interpretable descriptions of the aggregate decision rule. It is also much faster com-putationally making it more suitable to large scale data mining applications .
منابع مشابه
Additive Logistic Regression a Statistical View of Boosting
Boosting Freund Schapire Schapire Singer is one of the most important recent developments in classi cation method ology The performance of many classi cation algorithms often can be dramatically improved by sequentially applying them to reweighted versions of the input data and taking a weighted majority vote of the sequence of classi ers thereby produced We show that this seemingly mysterious ...
متن کاملMultiplicative Models in Projection Pursuit
Friedman and Stuetzle (JASA, 1981) developed a methodology for modeling a response surface by the sum of general smooth functions of linear combinations of the predictor variables. Here multiplicative models for regression and categorical regression are explored. The construction of these models and their performance relative to additive models are examined. CHAPTER 0 INTRODUCTION In recent wor...
متن کاملImproving Credit Scoring by Generalized Additive Model
Logistic Regression has been widely used in the financial service industry for credit scoring models. Despite its advantages in easy interpretation and low computing cost, Logistic Regression is under the criticism of failure to model the nonlinear features of the predictors effect on the dependent variable and therefore might lead to unsatisfactory results. Modern statistical techniques such a...
متن کاملHybrid Method of Logistic Regression and Data Envelopment Analysis for Event Prediction: A Case Study (Stroke Disease)
Abstract Predictive analytics is an area of statistics that deals with extracting information from data and using it to predict trends and behavior patterns. Many mathematical modeling has been developed and used for prediction, and in some cases, they have been found to be very strong and reliable. This paper studies different mathematical and statistical approaches for events prediction. The ...
متن کامل